Papers with Tatoeba database
TaPaCo: A Corpus of Sentential Paraphrases for 73 Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | a crowdsourcing project aimed at language learners has created a paraphrase corpus for 73 languages . the corpus contains 1.9 million sentences, with 200 - 250 000 sentences per language . |
| Approach: | They propose to use a Tatoeba-based dataset to create a paraphrase corpus for 73 languages. |
| Outcome: | The proposed dataset contains 1.9 million sentences and 200 - 250 000 sentences per language. |